Abstract
Background: Deep learning auto-segmentation is entering routine radiotherapy planning, and its performance is reported almost universally as a geometric overlap statistic. Whether that statistic establishes clinical readiness depends on how it relates to the reproducibility of the reference standard and to dose delivered. Methods: Systematic reviews and clinical evaluations of deep learning auto-segmentation in radiotherapy reporting evaluation methodology, overlap statistics or dosimetric outcomes were eligible. Reported metric usage was extracted. The fraction of a structure lying within a boundary shell of given thickness was derived analytically to quantify the sensitivity of volume-overlap statistics to boundary error. No pooling was performed. Analyses were performed in Python 3. Results: Across 117 reviewed studies, a geometric metric was used in 116 (99.1%) and the Dice similarity coefficient specifically in 113 (96.6%). Dosimetric assessment appeared in 27 (23.1%), qualitative or physician rating in 22 (18.8%), time saving in 18 (15.4%) and editing time in 11 (9.4%). Over 90 different names were used for geometric measures. A single manual contour served as ground truth in 65 studies (55.6%), and only 31 (26.5%) compared auto-contours with inter- or intra-observer variation, meaning 74% did not establish the reproducibility of their own reference standard. Reported expert reproducibility for a delineation task was a Dice of approximately 0.80 within observers and 0.67 between them; a deep learning model scored 0.76 against intra-observer 0.77 (p = 0.307) and 0.78 against inter-observer 0.67. In a commercial software evaluation with a mean Dice of 0.72, described as acceptable in clinical practice, the volume receiving 45 Gy in a rectal case was 10.4 cm³ by manual contour and 289.4 cm³ by automatic contour, a 27.8-fold difference. Conclusions. The metric that 97% of studies report is one the field states is not predictable of dosimetric impact, is measured against a reference standard whose own reproducibility is usually unestablished, and is structurally insensitive to boundary error because a 2mm shell constitutes under a fifth of the volume of a 30 mm structure. Scores exceeding inter-observer agreement cannot be read as accuracy. Evaluation should pair overlap statistics with dosimetric assessment and physician rating, both of which appear in fewer than a quarter of published studies.
Keywords: Auto-segmentation, Deep learning, Dice similarity coefficient, Inter-observer variability, Radiotherapy planning, Dosimetric evaluation